NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Mosaic Pages: Big TLB Reach With Small Pages

https://doi.org/10.1109/MM.2024.3409181

Han, Jaehyun; Gosakan, Krishnan; Kuszmaul, William; Mubarek, Ibrahim N; Mukherjee, Nirjhar; Sriram, Karthik; Tagliavini, Guido; West, Evan; Bender, Michael A; Bhattacharjee, Abhishek; et al (July 2024, IEEE Micro)

Full Text Available
Mosaic Pages: Big TLB Reach with Small Pages

https://doi.org/10.1145/3582016.3582021

Gosakan, Krishnan; Han, Jaehyun; Kuszmaul, William; Mubarek, Ibrahim N.; Mukherjee, Nirjhar; Sriram, Karthik; Tagliavini, Guido; West, Evan; Bender, Michael A.; Bhattarcharjee, Abhishek; et al (March 2023, ASPLOS)

Full Text Available
BetrFS: a compleat file system for commodity SSDs

https://doi.org/10.1145/3492321.3519571

Jiao, Yizheng; Bertron, Simon; Patel, Sagar; Zeller, Luke; Bennett, Rory; Mukherjee, Nirjhar; Bender, Michael A.; Condict, Michael; Conway, Alex; Farach-Colton, Martín; et al (March 2022, BetrFS: a compleat file system for commodity SSDs)

Full Text Available
External-memory Dictionaries in the Affine and PDAM Models

https://doi.org/10.1145/3470635

Bender, Michael A.; Conway, Alex; Farach-Colton, Martín; Jannen, William; Jiao, Yizheng; Johnson, Rob; Knorr, Eric; Mcallister, Sara; Mukherjee, Nirjhar; Pandey, Prashant; et al (September 2021, ACM Transactions on Parallel Computing)
null (Ed.)
Storage devices have complex performance profiles, including costs to initiate IOs (e.g., seek times in hard drives), parallelism and bank conflicts (in SSDs), costs to transfer data, and firmware-internal operations. The Disk-access Machine (DAM) model simplifies reality by assuming that storage devices transfer data in blocks of size B and that all transfers have unit cost. Despite its simplifications, the DAM model is reasonably accurate. In fact, if B is set to the half-bandwidth point, where the latency and bandwidth of the hardware are equal, then the DAM approximates the IO cost on any hardware to within a factor of 2. Furthermore, the DAM model explains the popularity of B-trees in the 1970s and the current popularity of B ɛ -trees and log-structured merge trees. But it fails to explain why some B-trees use small nodes, whereas all B ɛ -trees use large nodes. In a DAM, all IOs, and hence all nodes, are the same size. In this article, we show that the affine and PDAM models, which are small refinements of the DAM model, yield a surprisingly large improvement in predictability without sacrificing ease of use. We present benchmarks on a large collection of storage devices showing that the affine and PDAM models give good approximations of the performance characteristics of hard drives and SSDs, respectively. We show that the affine model explains node-size choices in B-trees and B ɛ -trees. Furthermore, the models predict that B-trees are highly sensitive to variations in the node size, whereas B ɛ -trees are much less sensitive. These predictions are born out empirically. Finally, we show that in both the affine and PDAM models, it pays to organize data structures to exploit varying IO size. In the affine model, B ɛ -trees can be optimized so that all operations are simultaneously optimal, even up to lower-order terms. In the PDAM model, B ɛ -trees (or B-trees) can be organized so that both sequential and concurrent workloads are handled efficiently. We conclude that the DAM model is useful as a first cut when designing or analyzing an algorithm or data structure but the affine and PDAM models enable the algorithm designer to optimize parameter choices and fill in design details.
more » « less
Full Text Available
Copy-on-Abundant-Write for Nimble File System Clones

Zhan, Yang; Conway, Alex; Jiao, Yizheng; Mukherjee, Nirjhar; Groombridge, Ian; Bender, Michael; Farach-Colton, Martin; Jannen, William; Johnson, Rob; Porter, Donald; et al (January 2021, ACM transactions on storage)

Making logical copies, or clones, of files and directories is critical to many real-world applications and work- flows, including backups, virtual machines, and containers. An ideal clone implementation meets the follow- ing performance goals: (1) creating the clone has low latency; (2) reads are fast in all versions (i.e., spatial locality is always maintained, even after modifications); (3) writes are fast in all versions; (4) the overall sys- tem is space efficient. Implementing a clone operation that realizes all four properties, which we call a nimble clone, is a long-standing open problem. This article describes nimble clones in B-ε-tree File System (BetrFS), an open-source, full-path-indexed, and write-optimized file system. The key observation behind our work is that standard copy-on-write heuristics can be too coarse to be space efficient, or too fine-grained to preserve locality. On the other hand, a write- optimized key-value store, such as a Bε -tree or an log-structured merge-tree (LSM)-tree, can decouple the logical application of updates from the granularity at which data is physically copied. In our write-optimized clone implementation, data sharing among clones is only broken when a clone has changed enough to warrant making a copy, a policy we call copy-on-abundant-write. We demonstrate that the algorithmic work needed to batch and amortize the cost of BetrFS clone operations does not erode the performance advantages of baseline BetrFS; BetrFS performance even improves in a few cases. BetrFS cloning is efficient; for example, when using the clone operation for container creation, BetrFSoutperforms a simple recursive copy by up to two orders-of-magnitude and outperforms file systems that have specialized Linux Containers (LXC) backends by 3–4×.
more » « less
Full Text Available
Copy-on-Abundant-Write for Nimble File System Clones

https://doi.org/10.1145/3423495

Zhan, Yang; Conway, Alex; Jiao, Yizheng; Mukherjee, Nirjhar; Groombridge, Ian; Bender, Michael A.; Farach-Colton, Martin; Jannen, William; Johnson, Rob; Porter, Donald E.; et al (February 2021, ACM Transactions on Storage)
null (Ed.)
Making logical copies, or clones, of files and directories is critical to many real-world applications and workflows, including backups, virtual machines, and containers. An ideal clone implementation meets the following performance goals: (1) creating the clone has low latency; (2) reads are fast in all versions (i.e., spatial locality is always maintained, even after modifications); (3) writes are fast in all versions; (4) the overall system is space efficient. Implementing a clone operation that realizes all four properties, which we call a nimble clone , is a long-standing open problem. This article describes nimble clones in B-ϵ-tree File System (BetrFS), an open-source, full-path-indexed, and write-optimized file system. The key observation behind our work is that standard copy-on-write heuristics can be too coarse to be space efficient, or too fine-grained to preserve locality. On the other hand, a write-optimized key-value store, such as a Bε-tree or an log-structured merge-tree (LSM)-tree, can decouple the logical application of updates from the granularity at which data is physically copied. In our write-optimized clone implementation, data sharing among clones is only broken when a clone has changed enough to warrant making a copy, a policy we call copy-on-abundant-write . We demonstrate that the algorithmic work needed to batch and amortize the cost of BetrFS clone operations does not erode the performance advantages of baseline BetrFS; BetrFS performance even improves in a few cases. BetrFS cloning is efficient; for example, when using the clone operation for container creation, BetrFS outperforms a simple recursive copy by up to two orders-of-magnitude and outperforms file systems that have specialized Linux Containers (LXC) backends by 3--4×.
more » « less
Full Text Available
Paging and the Address-Translation Problem

https://doi.org/10.1145/3409964.3461814

Bender, Michael A.; Bhattacharjee, Abhishek; Conway, Alex; Farach-Colton, Martín; Johnson, Rob; Kannan, Sudarsun; Kuszmaul, William; Mukherjee, Nirjhar; Porter, Don; Tagliavini, Guido; et al (January 2021, 33rd ACM Symposium on Parallelism in Algorithms and Architectures (SPAA))
null (Ed.)
Full Text Available
External-Memory Dictionaries in the Affine and PDAM Models

https://doi.org/https://doi.org/10.1145/3323165.3323210

Bender, Michael; Conway, Alex; Farach-Colton, Martin; Jannen, William; Jiao, Yizheng; Johnson, Rob; Knorr, Eric; McAllister, Sara; Mukherjee, Nirjhar; Pandey, Prashant; et al (January 2021, ACM transactions on parallel computing)

Storage devices have complex performance profiles, including costs to initiate IOs (e.g., seek times in hard 15 drives), parallelism and bank conflicts (in SSDs), costs to transfer data, and firmware-internal operations. The Disk-access Machine (DAM) model simplifies reality by assuming that storage devices transfer data in blocks of size B and that all transfers have unit cost. Despite its simplifications, the DAM model is reasonably accurate. In fact, if B is set to the half-bandwidth point, where the latency and bandwidth of the hardware are equal, then the DAM approximates the IO cost on any hardware to within a factor of 2. Furthermore, the DAM model explains the popularity of B-trees in the 1970s and the current popularity of Bε -trees and log-structured merge trees. But it fails to explain why some B-trees use small nodes, whereas all Bε -trees use large nodes. In a DAM, all IOs, and hence all nodes, are the same size. In this article, we show that the affine and PDAM models, which are small refinements of the DAM model, yield a surprisingly large improvement in predictability without sacrificing ease of use. We present benchmarks on a large collection of storage devices showing that the affine and PDAM models give good approximations of the performance characteristics of hard drives and SSDs, respectively. We show that the affine model explains node-size choices in B-trees and Bε -trees. Furthermore, the models predict that B-trees are highly sensitive to variations in the node size, whereas Bε -trees are much less sensitive. These predictions are born out empirically. Finally, we show that in both the affine and PDAM models, it pays to organize data structures to exploit varying IO size. In the affine model, Bε -trees can be optimized so that all operations are simultaneously optimal, even up to lower-order terms. In the PDAM model, Bε -trees (or B-trees) can be organized so that both sequential and concurrent workloads are handled efficiently. We conclude that the DAM model is useful as a first cut when designing or analyzing an algorithm or data structure but the affine and PDAM models enable the algorithm designer to optimize parameter choices and fill in design details.
more » « less
Full Text Available

Search for: All records